Papers with scene-level segmentation
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding (2025.emnlp-industry)
Copied to clipboard
Xiangfeng Wang, Xiao Li, Yadong Wei, null Songxueyu, Yang Song, null Xiaxiaoqiang, Fangrui Zeng, Zaiyi Chen, null Liuliu, Gu Xu, Tong Xu
| Challenge: | Existing methods for video editing rely on textual cues from ASR transcripts and segment selection, often neglecting rich visual context. |
| Approach: | They propose a human-inspired automatic video editing framework that leverages multimodal narrative understanding to address these limitations. |
| Outcome: | The proposed framework outperforms existing baselines across general and advertisement-oriented editing tasks. |